Papers with reward prediction
Large Language Model-Enhanced Multi-Armed Bandits (2026.acl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) have been used to sequential decision-making tasks like multi-armed bandits where an LLM is tasked with selecting arms in each iteration is often suboptimal. |
| Approach: | They propose to combine MAB and LLMs to leverage the in-context learning capability of LLM for reward prediction. |
| Outcome: | The proposed approach outperforms LLM-based direct arm selection on synthetic tasks where only preference feedback between arm pairs is available. |
Self-Generated Critiques Boost Reward Modeling for Language Models (2025.naacl-long)
Copied to clipboard
Yue Yu, Zhengxing Chen, Aston Zhang, Liang Tan, Chenguang Zhu, Richard Yuanzhe Pang, Yundi Qian, Xuewei Wang, Suchin Gururangan, Chao Zhang, Melanie Kambadur, Dhruv Mahajan, Rui Hou
| Challenge: | Existing reward models produce scalar scores and struggle to incorporate critiques in a natural language format. |
| Approach: | They propose a framework that predicts critiques and rewards using self-generated critiques without extra supervision. |
| Outcome: | The proposed framework improves reward modeling accuracy by 3.7%-7.3% compared to standard reward models and LLM judges. |
P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic Checklist (2026.acl-long)
Copied to clipboard
| Challenge: | Existing approaches to personalized reward modeling treat user context as static or implicit conditioning signal, failing to capture dynamic nature of human judgment. |
| Approach: | They propose a personalized reward modeling framework that synthesizes dynamic evaluation criteria for guiding the reward prediction. |
| Outcome: | The proposed framework improves reward accuracy and enhances downstream personalized generation. |